Papers with cognitive tests
MovieCORE: COgnitive REasoning in Movies (2025.emnlp-main)
Copied to clipboard
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh, Ying Cheng, Hung-Ting Su, Yung-Hao Tang, Shang-Hong Lai, Winston H. Hsu
| Challenge: | MovieCORE is a video question answering dataset that focuses on surface-level comprehension. |
| Approach: | They propose a video question-answer dataset that uses large language models as thought agents to generate and refine high-quality question-anchor pairs. |
| Outcome: | The proposed model improves model reasoning capabilities post-training by 25% . the proposed model is based on a large language model and is scalable to a wide range of tasks . |
Detecting Dementia from Long Neuropsychological Interviews (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies suggest examiner's language can influence cognitive impairment classifications. |
| Approach: | They propose a three-stage pipeline to detect dementia from exam recordings to mitigate the influence of the examiner on automatic dementia identification decisions. |
| Outcome: | The proposed pipeline mitigates the influence of the examiner on automatic dementia identification decisions in real-world neuropsychological exams. |
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests (2025.findings-emnlp)
Copied to clipboard
Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi
| Challenge: | MMLU and BBH are three evaluation paradigms for language learning models . interactive games are superior to standard benchmarks in discriminating models based on human cognitive assessments . |
| Approach: | They examine three evaluation paradigms: standard benchmarks, interactive games and cognitive tests . they examine whether interactive games are more effective at discriminating LLMs . |
| Outcome: | The results show that interactive games are superior to standard benchmarks in discriminating models. |